Supersonic La2026-09-26 19:54:30Supersonic Labs releases Julia 1, a 144.3 million-parameter open-source decision model that runs on CPUsBrazilian AI lab Supersonic Labs has released Julia 1, an open-source decision model built for structured decision tasks rather than general-purpose language generation. The model has 144.3 million parameters and is designed to run locally on standard CPUs, removing the need for high-end GPUs. Its weights have been published on Hugging Face under the Apache 2.0 license, and the model also supports browser-side execution through ONNX with WebGPU. Julia 1 uses a single API to handle three types of decision workflows: choosing from 2 to 20 options, assigning ratings on ordered scales such as low, medium, and high, and estimating the probability that a yes-or-no statement is true. Supersonic Labs said the model is built on Johns Hopkins University’s mmBERT-small encoder and is not a fine-tuned version of an existing large language model. The lab put total training cost at about $104. In benchmark results released by the team, Julia 1 posted 73.15% accuracy on the Typed Decisions task, slightly above the reference baseline for the TypeSafe Jev model. Performance was weaker on some other tasks, including Banking77, where accuracy reached 64% across a 72-label classification setting. On Apple’s M4 chip, the median latency for a single decision was 33.15 milliseconds.20
Kimi K32026-08-08 06:24:422.78T-Parameter Kimi K3 Runs on 8GB RAM via Open-Source C99 CPU EngineAn open-source project called kimi-k3-in-c aims to run Kimi K3, a model with 2.78 trillion parameters, on devices with just 8GB of memory. The 176KB codebase is written in pure C99 and performs inference on the CPU only, dropping GPU, CUDA, PyTorch, and BLAS entirely. The approach exploits Kimi K3's MoE architecture: only 16 of the 896 experts per layer are activated, so the developer avoids loading the full ~1.56TB of weights and instead streams most expert weights from NVMe storage on demand. Dense trunk layers are streamed layer by layer as well. Trade-offs remain: generating one token takes about 32.7 seconds, and close to 1.7TB of fast storage is needed. The developer describes the project as an experimental exploration of LLM inference infrastructure rather than a production-ready solution, but the combination of hard-drive streaming and sparse MoE activation points toward new ways to run ultra-large models at low cost.2280